Repository navigation
fix: bound ShellCheck concurrency across fm-lint runs - #71
Merged
Merged
Conversation
Each fm-lint.sh run limited itself to two workers, but nothing limited the sum, so N concurrent runs started 2N ShellCheck processes (42 seen, some near 4.8 GB). Every root now holds one host-wide flock slot: half the cores (at least two), shrinking to a floor of two as 1-minute load passes twice the cores. Runs queue instead of failing, and queue time stays out of the root deadline and recorded duration.
Bound every waiting path to one fresh load probe per two seconds while preserving occupancy revalidation. Force two fixture workers and hold checks until the controller observes the required admission count. Replace scheduler-sensitive timestamp comparisons with observed lifecycle brackets. Counted architecture comparison (binary factors): Design | Host-wide bound | Preserves two workers | Resident daemon | Custom stale-lock recovery | Extra per-root gate process Per-run only | 0 | 1 | 0 | 0 | 0 FM_LINT_JOBS=1 | 0 | 0 | 0 | 0 | 0 Bash mkdir slots | 1 | 1 | 0 | 1 | 0 Central daemon, disconnect-based release | 1 | 1 | 1 | 0 | 0 Flock pool | 1 | 1 | 0 | 0 | 1 The flock pool supplies the required host-wide bound and preserves two workers without a resident service or custom stale-lock recovery, at the cost of one gate process per root. Paired reproduction: six simultaneous fm-lint runs, four identical roots per run, --jobs 2, inherited FM_LINT_JOBS=1, and controller-held fake ShellCheck processes. With FM_LINT_SLOT_DIR=off, the measured peak was 12 = 2 x 6 runs. With a three-slot pool and zero simulated load, the measured peak was 3 = the slot cap. Focused verification: all seven host-slot regressions passed, covering concurrency, load thresholds and occupancy, disable/misconfiguration behavior, slot I/O and locking failures, gate-death ownership, and queue-excluded initial/retry timing. A real load-reader smoke observed two probes over 4.5 seconds on both the free-capacity and blocking-wakeup paths, with intervals of 2.622 and 2.653 seconds respectively; admission resumed below the occupancy allowance. No full repository test or lint suite was run. The outer executor owns copying this counted comparison and paired reproduction into the PR description. No new tracked documentation was added.
…ency documentation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
I would like all bugs to be fixed so tomorrow can be focused entirely to Vernant and not problems preventing Vernant from getting built. A recommended architecture should come with proof.
Context: the backlog item "Bound fm-lint shellcheck fan-out" asks to limit how many shellcheck processes fm-lint.sh runs at once by hw.ncpu and current load, queueing instead of failing (up to 42 ran at once, some about 4.8 GB each). It is not a cap on agents. Source: a host analysis of a 21-minute sample at load 102 on 18 cores: 0.1% idle, 56% kernel time, about 61% of CPU in sub-second processes, about 2,850 new processes per second. The target for the whole set of fixes over one fleet day is 1-minute load <= 36 for >= 90% of the day, process creation < 500/s, kernel time < 25%, and no new timeout-class failures. Pattern: measure, fix the root, test, measure again.
Decisions made in this change: the root cause is that fm-lint.sh bounds concurrency per run only (FM_LINT_JOBS), so N concurrent runs start 2N ShellCheck processes; a single-run test hides it. The fix is a host-wide counting semaphore of flock slot files taken per root by a gate mode in bin/fm-lint-cache.pl (Perl flock works on macOS, which has no flock command, and the kernel frees a slot when its holder dies, so no stale-lock code). Slot count is FM_LINT_HOST_SLOTS else max(2, ncpu/2); the allowance shrinks as 1-minute load rises past 2 x ncpu (matches the load target), with a floor of 2 so a lone run is never throttled and queued roots always progress. Waiting runs queue rather than fail. The gate sits outside the per-root timeout watchdog and queue time is excluded from the root deadline and recorded duration, so no new timeout-class failures. An unusable slot directory runs ungated with a warning; FM_LINT_SLOT_DIR=off disables the gate. A bad FM_LINT_HOST_SLOTS exits 2. Proof: five designs were compared on counted factors (per-run only, FM_LINT_JOBS=1, bash mkdir slots, central daemon, flock pool), and a fake-ShellCheck reproduction shows peak live processes of 2 x runs before and the slot cap after. The new behavioral tests in tests/fm-lint.test.sh fail on the pre-fix code and pass after. Not in scope: fleet-day measurement, FIFO ordering of waiters, queue-wait telemetry.
What Changed
flockslot pool that bounds ShellCheck concurrency across lint runs using CPU count and current load, with inherited locks protecting running descendants.Risk Assessment
✅ Low: The change is narrowly scoped to host-wide lint admission and its behavioral proof, with no substantiated correctness, privacy, or intent-conformance defects found in the current source.
Testing
Targeted CLI and gate scenarios passed with real ShellCheck, including counted contention, load admission, queue timing, diagnostics, and process cleanup. The base-commit concurrency regression reproduced the original failure. Scheduler-sensitive test-controller waits were fixed and exercised afterward. CLI transcripts, process snapshots, telemetry, regression results, and native capability limits were retained; live memory-exhaustion retry could not be exercised. No repository-wide suite or separate lint/static-analysis phase was run.
Evidence: Real-product concurrency, load, queue, and lifetime evidence
Source: Real-product concurrency, load, queue, and lifetime evidence
Evidence: CLI configuration guards and diagnostic parity
Source: CLI configuration guards and diagnostic parity
Evidence: Targeted regressions and failing-before concurrency proof
Source: Targeted regressions and failing-before concurrency proof
Evidence: Native memory-bound and heap-limit capability blockers
Source: Native memory-bound and heap-limit capability blockers
Evidence: Actual ShellCheck contention instrumentation
Source: Actual ShellCheck contention instrumentation
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 4 issues found → auto-fixed (4) → no changes applied ✅
bin/fm-lint-cache.pl:60- The allowance contradicts the required criterion: "the allowance shrinks as 1-minute load rises past 2 x ncpu." The added calculation is int(2 * ncpu - load + 0.5), clamped to the cap, rather than subtracting only excess load from the cap. With 18 cores and the default cap of 9, load 30 permits only 6 slots and load 36 permits only 2; the stated behavior requires all 9 until load exceeds 36. The same contract is documented at bin/fm-lint-cache.pl:21 and bin/fm-lint.sh:68. Tests at tests/fm-lint.test.sh:1165 and :1169 exercise only extreme and zero load, missing this boundary. Reconcile the implementation with the authorized threshold by reducing the cap only for load above twice the core count.bin/fm-lint-cache.pl:64- Slot-file failures are treated as contention and can queue a root forever. For example, a writable slot directory whose existing allowed slot files are non-writable passes the directory check at :56, but every open at :64 and :71 fails; the loop sleeps and retries indefinitely without warning or starting analysis. Exhausted filesystem metadata produces the same path. Non-contention flock failures at :65 and :75 can instead produce a tight retry loop. Because the gate is outside the watchdog at bin/fm-lint.sh:297, no root deadline terminates either case. Distinguish busy locks from slot-pool I/O/locking failures at the shared gate boundary and use the already-authorized warning-and-ungated fallback at :85.bin/fm-lint-cache.pl:93- The slot can be released while its bounded ShellCheck tree is still running. Perl normally closes the opened slot descriptor on exec, so after the fork/exec at :93-96 only the gate retains the lock acquired at :65 or :80. On ordinary lint cancellation, bin/fm-lint.sh:1088 immediately TERM/KILLs the worker group, including this new gate. The watchdog launched through bin/fm-lint.sh:297 is in a separate group, and its analysis tree remains alive until its owner-death polling and TERM/KILL cleanup complete (bin/fm-timeout-lib.sh:318-336). An already-queued gate can therefore acquire the released slot and launch another analysis before the previous one exits, exceeding the advertised host-wide bound. Preserve the same slot lock through the protected command/watchdog lifetime, including gate death; ensure the launch at bin/fm-lint.sh:306 follows the same ownership invariant.bin/fm-lint.sh:394- Subtracting millisecond queue time from second-resolution timestamps can record negative root durations on Bash versions without EPOCHREALTIME, including macOS's system Bash. fm_lint_now_ms at :170 then truncates timestamps to whole seconds, while bin/fm-lint-cache.pl:89 records precise wait milliseconds. A root starting at T+0.1s, queueing for 1.5s, and finishing analysis at T+1.7s produces start=T000, end=(T+1)000, and duration_ms=-500 despite successful execution. The same mixed-precision inputs affect the retry budget at bin/fm-lint.sh:353; wait values enter through :313-315 and accumulate at :346 and :373. Use a consistent high-resolution clock for root timestamps and queue accounting rather than merely clamping the resulting duration.🔧 Fix applied.
3 issues (2 errors, 1 warning) still open:
bin/fm-lint-cache.pl:64- Slot-file failures are treated as contention and can queue a root forever. For example, a writable slot directory whose existing allowed slot files are non-writable passes the directory check at :56, but every open at :64 and :71 fails; the loop sleeps and retries indefinitely without warning or starting analysis. Exhausted filesystem metadata produces the same path. Non-contention flock failures at :65 and :75 can instead produce a tight retry loop. Because the gate is outside the watchdog at bin/fm-lint.sh:297, no root deadline terminates either case. Distinguish busy locks from slot-pool I/O/locking failures at the shared gate boundary and use the already-authorized warning-and-ungated fallback at :85.bin/fm-lint-cache.pl:93- The slot can be released while its bounded ShellCheck tree is still running. Perl normally closes the opened slot descriptor on exec, so after the fork/exec at :93-96 only the gate retains the lock acquired at :65 or :80. On ordinary lint cancellation, bin/fm-lint.sh:1088 immediately TERM/KILLs the worker group, including this new gate. The watchdog launched through bin/fm-lint.sh:297 is in a separate group, and its analysis tree remains alive until its owner-death polling and TERM/KILL cleanup complete (bin/fm-timeout-lib.sh:318-336). An already-queued gate can therefore acquire the released slot and launch another analysis before the previous one exits, exceeding the advertised host-wide bound. Preserve the same slot lock through the protected command/watchdog lifetime, including gate death; ensure the launch at bin/fm-lint.sh:306 follows the same ownership invariant.bin/fm-lint.sh:394- Subtracting millisecond queue time from second-resolution timestamps can record negative root durations on Bash versions without EPOCHREALTIME, including macOS's system Bash. fm_lint_now_ms at :170 then truncates timestamps to whole seconds, while bin/fm-lint-cache.pl:89 records precise wait milliseconds. A root starting at T+0.1s, queueing for 1.5s, and finishing analysis at T+1.7s produces start=T000, end=(T+1)000, and duration_ms=-500 despite successful execution. The same mixed-precision inputs affect the retry budget at bin/fm-lint.sh:353; wait values enter through :313-315 and accumulate at :346 and :373. Use a consistent high-resolution clock for root timestamps and queue accounting rather than merely clamping the resulting duration.🔧 Fix applied.
6 issues (5 errors, 1 warning) still open:
bin/fm-lint-cache.pl:64- Slot-file failures are treated as contention and can queue a root forever. For example, a writable slot directory whose existing allowed slot files are non-writable passes the directory check at :56, but every open at :64 and :71 fails; the loop sleeps and retries indefinitely without warning or starting analysis. Exhausted filesystem metadata produces the same path. Non-contention flock failures at :65 and :75 can instead produce a tight retry loop. Because the gate is outside the watchdog at bin/fm-lint.sh:297, no root deadline terminates either case. Distinguish busy locks from slot-pool I/O/locking failures at the shared gate boundary and use the already-authorized warning-and-ungated fallback at :85.bin/fm-lint-cache.pl:93- The slot can be released while its bounded ShellCheck tree is still running. Perl normally closes the opened slot descriptor on exec, so after the fork/exec at :93-96 only the gate retains the lock acquired at :65 or :80. On ordinary lint cancellation, bin/fm-lint.sh:1088 immediately TERM/KILLs the worker group, including this new gate. The watchdog launched through bin/fm-lint.sh:297 is in a separate group, and its analysis tree remains alive until its owner-death polling and TERM/KILL cleanup complete (bin/fm-timeout-lib.sh:318-336). An already-queued gate can therefore acquire the released slot and launch another analysis before the previous one exits, exceeding the advertised host-wide bound. Preserve the same slot lock through the protected command/watchdog lifetime, including gate death; ensure the launch at bin/fm-lint.sh:306 follows the same ownership invariant.bin/fm-lint-cache.pl:67- Round 1 corrected the load calculation but left admission based on slot indices rather than total occupancy. Concrete sequence: on 18 cores, nine checks acquire slots at low load; load then rises to 43, making the allowance two. When slot.0 finishes, this branch immediately admits its replacement even while slots.1–8 remain occupied, restoring nine live checks despite the two-check allowance. The blocking acquisition at bin/fm-lint-cache.pl:75–87 also admits a slot selected before the allowance shrank without revalidating it. At the shared gate boundary, let existing checks finish but queue new admissions until total occupied capacity is below the current allowance; apply that invariant to both acquisition paths.tests/fm-lint.test.sh:1398- Round 2 introduced a lifetime regression that rejects correct watchdog cleanup on bounded hosts. Killing the gate changes the watchdog's parent; bin/fm-timeout-lib.sh:332–334 consequently terminates the protected group. If its time/cache wrapper exits, :312–316 kills the remaining group, legitimately releasing the inherited slot. Nevertheless, this assertion forbids admission after a fixed 200 ms without establishing that the old tree remains alive. The same incorrect assumption appears at tests/fm-lint.test.sh:1401, and :1402 requires a descendant to survive cleanup that the watchdog intentionally performs. [INFERENCE from source] These assertions can fail with correct slot ownership. Keep the survivor checks in the unbounded fixture; in the bounded fixture, assert that admission does not overlap a live protected tree and permit admission after cleanup.docs/fm-test-portable-shards.md:139- The required architecture proof is absent from the reviewed deliverable. Intent specifies: "five designs were compared on counted factors" and a reproduction showing "peak live processes of 2 x runs before and the slot cap after." The added documentation at :139–142 describes implementation and test invocation but supplies no counted comparison. The concurrency test at tests/fm-lint.test.sh:1148–1156 measures only gated runs; the disabled-gate test at :1243 executes one root without measuring concurrency, so neither supplies the paired baseline. Retain the five-design comparison and a counted ungated/gated reproduction, or obtain explicit approval to waive those proof requirements. This is missing source-verifiable evidence, not a request to execute the pipeline's later validation phase.tests/fm-lint.test.sh:3188- Simplification: Round 2 added the --host-slots test mode, duplicating the seven-test dispatch already present at tests/fm-lint.test.sh:3236–3242. The authorized findings require behavioral regressions, not another test-runner mode; the existing suite satisfies that requirement. Revert this component to the minimal fix by removing the additional dispatch branch and its documentation at docs/fm-test-portable-shards.md:142, while retaining the regression tests in the normal suite.🔧 Fix applied.
8 issues (5 errors, 3 warnings) still open:
bin/fm-lint-cache.pl:64- Slot-file failures are treated as contention and can queue a root forever. For example, a writable slot directory whose existing allowed slot files are non-writable passes the directory check at :56, but every open at :64 and :71 fails; the loop sleeps and retries indefinitely without warning or starting analysis. Exhausted filesystem metadata produces the same path. Non-contention flock failures at :65 and :75 can instead produce a tight retry loop. Because the gate is outside the watchdog at bin/fm-lint.sh:297, no root deadline terminates either case. Distinguish busy locks from slot-pool I/O/locking failures at the shared gate boundary and use the already-authorized warning-and-ungated fallback at :85.bin/fm-lint-cache.pl:93- The slot can be released while its bounded ShellCheck tree is still running. Perl normally closes the opened slot descriptor on exec, so after the fork/exec at :93-96 only the gate retains the lock acquired at :65 or :80. On ordinary lint cancellation, bin/fm-lint.sh:1088 immediately TERM/KILLs the worker group, including this new gate. The watchdog launched through bin/fm-lint.sh:297 is in a separate group, and its analysis tree remains alive until its owner-death polling and TERM/KILL cleanup complete (bin/fm-timeout-lib.sh:318-336). An already-queued gate can therefore acquire the released slot and launch another analysis before the previous one exits, exceeding the advertised host-wide bound. Preserve the same slot lock through the protected command/watchdog lifetime, including gate death; ensure the launch at bin/fm-lint.sh:306 follows the same ownership invariant.bin/fm-lint-cache.pl:67- Round 1 corrected the load calculation but left admission based on slot indices rather than total occupancy. Concrete sequence: on 18 cores, nine checks acquire slots at low load; load then rises to 43, making the allowance two. When slot.0 finishes, this branch immediately admits its replacement even while slots.1–8 remain occupied, restoring nine live checks despite the two-check allowance. The blocking acquisition at bin/fm-lint-cache.pl:75–87 also admits a slot selected before the allowance shrank without revalidating it. At the shared gate boundary, let existing checks finish but queue new admissions until total occupied capacity is below the current allowance; apply that invariant to both acquisition paths.tests/fm-lint.test.sh:1398- Round 2 introduced a lifetime regression that rejects correct watchdog cleanup on bounded hosts. Killing the gate changes the watchdog's parent; bin/fm-timeout-lib.sh:332–334 consequently terminates the protected group. If its time/cache wrapper exits, :312–316 kills the remaining group, legitimately releasing the inherited slot. Nevertheless, this assertion forbids admission after a fixed 200 ms without establishing that the old tree remains alive. The same incorrect assumption appears at tests/fm-lint.test.sh:1401, and :1402 requires a descendant to survive cleanup that the watchdog intentionally performs. [INFERENCE from source] These assertions can fail with correct slot ownership. Keep the survivor checks in the unbounded fixture; in the bounded fixture, assert that admission does not overlap a live protected tree and permit admission after cleanup.tests/fm-lint.test.sh:3188- Simplification: Round 2 added the --host-slots test mode, duplicating the seven-test dispatch already present at tests/fm-lint.test.sh:3236–3242. The authorized findings require behavioral regressions, not another test-runner mode; the existing suite satisfies that requirement. Revert this component to the minimal fix by removing the additional dispatch branch and its documentation at docs/fm-test-portable-shards.md:142, while retaining the regression tests in the normal suite.tests/fm-lint.test.sh:1152- The authorized proof requirement remains incomplete after the fix following Round 3. The recorded decision requires: "Put the five-design counted comparison and a paired reproduction in the PR description and in the commit message." Every commit message from the base through HEAD was inspected; none contains the counted five-design comparison or paired reproduction results. Lines 1152–1160 now provide an executable baseline/gated fixture, but that is not the required retained comparison and measured proof in the commit message. Supply that evidence in the commit message without adding tracked docs, or obtain an explicit waiver. The later PR update remains pipeline-owned and is not the subject of this finding.bin/fm-lint-cache.pl:79- The occupancy fix following Round 3 introduces rapid process creation precisely when host load is excessive. On an 18-core Mac at load 43, the nine-slot pool permits two occupants but leaves seven files unlocked. Waiting gates therefore repeatedly take the has-free branch at lines 79–81 and invoke the external sysctl load reader at lines 48–50 on each 100 ms retry. With the intended 21 concurrent runs, up to 40 roots can be waiting; this path can generate hundreds of short-lived sysctl processes per second despite starting no analysis. Preserve occupancy revalidation, but use a bounded, slower load-recheck cadence across this branch and the blocking-wait path at lines 83–95 rather than multiplying 10 Hz load probes by the waiter count.tests/fm-lint.test.sh:1123- The concurrent-count fixture does not establish the concurrency its assertions require. It inherits FM_LINT_JOBS, so running the suite with the documented FM_LINT_JOBS=1 immediately makes the lone-run assertion at line 1150 reject correct behavior. Independently, each stub exits after only 500 ms at line 1108: if a worker or later run is delayed beyond that window on the overloaded host this change targets, correct code produces a lower peak. The exact twelve-process baseline added after Round 3 at line 1153 inherits this problem, as do the gated saturation assertion at line 1159 and idle-load assertion at line 1174. Set --jobs 2 explicitly in the shared launcher and use controller-coordinated hold/release with bounded waits so the peak comparison measures admission rather than scheduler speed.🔧 No changes applied.
5 issues (2 errors, 3 warnings) still open:
tests/fm-lint.test.sh:1398- Round 2 introduced a lifetime regression that rejects correct watchdog cleanup on bounded hosts. Killing the gate changes the watchdog's parent; bin/fm-timeout-lib.sh:332–334 consequently terminates the protected group. If its time/cache wrapper exits, :312–316 kills the remaining group, legitimately releasing the inherited slot. Nevertheless, this assertion forbids admission after a fixed 200 ms without establishing that the old tree remains alive. The same incorrect assumption appears at tests/fm-lint.test.sh:1401, and :1402 requires a descendant to survive cleanup that the watchdog intentionally performs. [INFERENCE from source] These assertions can fail with correct slot ownership. Keep the survivor checks in the unbounded fixture; in the bounded fixture, assert that admission does not overlap a live protected tree and permit admission after cleanup.tests/fm-lint.test.sh:3188- Simplification: Round 2 added the --host-slots test mode, duplicating the seven-test dispatch already present at tests/fm-lint.test.sh:3236–3242. The authorized findings require behavioral regressions, not another test-runner mode; the existing suite satisfies that requirement. Revert this component to the minimal fix by removing the additional dispatch branch and its documentation at docs/fm-test-portable-shards.md:142, while retaining the regression tests in the normal suite.tests/fm-lint.test.sh:1152- The authorized proof requirement remains incomplete after the fix following Round 3. The recorded decision requires: "Put the five-design counted comparison and a paired reproduction in the PR description and in the commit message." Every commit message from the base through HEAD was inspected; none contains the counted five-design comparison or paired reproduction results. Lines 1152–1160 now provide an executable baseline/gated fixture, but that is not the required retained comparison and measured proof in the commit message. Supply that evidence in the commit message without adding tracked docs, or obtain an explicit waiver. The later PR update remains pipeline-owned and is not the subject of this finding.tests/fm-lint.test.sh:1123- The concurrent-count fixture does not establish the concurrency its assertions require. It inherits FM_LINT_JOBS, so running the suite with the documented FM_LINT_JOBS=1 immediately makes the lone-run assertion at line 1150 reject correct behavior. Independently, each stub exits after only 500 ms at line 1108: if a worker or later run is delayed beyond that window on the overloaded host this change targets, correct code produces a lower peak. The exact twelve-process baseline added after Round 3 at line 1153 inherits this problem, as do the gated saturation assertion at line 1159 and idle-load assertion at line 1174. Set --jobs 2 explicitly in the shared launcher and use controller-coordinated hold/release with bounded waits so the peak comparison measures admission rather than scheduler speed.tests/fm-lint.test.sh:1144- Round 4's controller releases every held check as soon as the expected cap is reached, before establishing that the remaining workers have attempted admission. With staggered startup, even the original per-run-only implementation can pass: three checks start, the controller releases them at :1149, and delayed workers subsequently finish without raising the peak above three. The ungated baseline waits for twelve, so the paired runs do not establish equivalent contention. The same early-release path affects the three-slot assertion at tests/fm-lint.test.sh:1181–1184 and heavy-load floor assertion at tests/fm-lint.test.sh:1193–1195; the idle-cap comparison at tests/fm-lint.test.sh:1197–1199 also uses this helper. Coordinate contender readiness and retain the admitted checks through a bounded contention-observation phase before releasing them, so removing host-wide admission reliably fails the regression.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
shellcheck --version: confirmed the available analyzer is ShellCheck 0.11.0.Drovebin/fm-lint.sh --jobs 2 --telemetry <isolated path> <six disposable roots>with one and four concurrent runs. Instrumentation suspended actual ShellCheck processes after exec, established contention, then resumed their real analyses.Compared four identical concurrent workloads withFM_LINT_SLOT_DIR=offversusFM_LINT_HOST_SLOTS=3: observed peaks of eight and three real ShellCheck processes.Executedperl bin/fm-lint-cache.pl gate ...with real ShellCheck and synthetic load inputs: default cap admitted nine processes at load 36 and eight at load 37; extreme load reduced an explicit six-slot pool to two.Exercised scan and lock-wakeup admission paths with eight occupied slot files and a load allowance of two; admission resumed only after occupancy fell below the allowance.Measured nativesysctlload probes while a gate was blocked; observed probe spacing exceeded two seconds.Drove the public CLI with its only slot held for five seconds; inspected generated root lifecycle and telemetry output to verify queue exclusion from analysis duration.Executed the actual gate andfm-lint.sh --internal-timedwatchdog with a six-second queue preceding a three-second analysis deadline; real ShellCheck completed successfully.Killed a gate and its immediate wrapper while a real ShellCheck descendant survived; verified capacity remained locked until that descendant exited.Killed a gate protecting the actual watchdog; verified contender admission followed cleanup without overlapping a live protected process.Drove invalid-cap, unusable-directory, unusable-slot-file, and real SC2086 finding scenarios through the public CLI; checked warnings, exit statuses, and gated/ungated diagnostic parity.Ran seven selected host-slot regression functions fromtests/fm-lint.test.sh, not the complete suite. Corrected controller waits and exercised the affected concurrency, I/O fallback, and timing checks afterward.Rantest_host_slots_bound_concurrent_runsagainst disposable executable copies from the supplied base commit; it failed with twelve processes against the expected three-slot bound.Attempted live memory-retry validation using native required bounds,GHCRTS=-M16m, and a read-only Docker availability probe; recorded the capability blockers.Removed all disposable workspace fixtures and temporary test drivers, and verified no owned live fixture processes remained.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.